Papers with Error analysis

17 papers
LaVA – Latvian Language Learner corpus (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of 1015 essays from foreigners learning Latvian as a foreign language is available at http://www.korpuss.lv/id/LaVA.
Approach: They propose to create a Latvian Language Learner Corpus (LaVA) which contains 1015 essays from Latvian students with different language backgrounds.
Outcome: The LaVA corpus contains 1015 essays from foreigners studying at Latvian higher education institutions and reaching the A1 (possibly A2) Latvian language proficiency level.
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs’ Memorization (2025.acl-long)

Copied to clipboard

Challenge: UnSeenTimeQA is a data contamination-free time-sensitive question-answering benchmark.
Approach: They propose a data contamination-free time-sensitive question-answering benchmark that avoids web-searchable queries grounded in the real world.
Outcome: The proposed benchmark avoids web-searchable queries grounded in the real world and enables on-demand generation of new samples, mitigating the risk of data leakage.
Can AMR Assist Legal and Logical Reasoning? (2022.findings-emnlp)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) has been shown to be useful for many downstream tasks.
Approach: They propose neural architectures that utilize linearised AMR graphs in combination with pre-trained language models to capture logical relationships on multiple choice question answering tasks.
Outcome: The proposed models outperform text-only baselines but outperformed text models, suggesting complementary abilities.
CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and Use (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks focus on narrow tasks such as multiple-choice cloze tests, isolated translation, or simple paraphrasing.
Approach: They propose a benchmark to measure Chinese idioms' cultural and contextual nuances . they evaluate 2,937 human-verified examples covering 1,765 common idiomes .
Outcome: The proposed benchmarks achieve 95% accuracy on Evaluative Connotation, but only 85% on Appropriateness and 40% top-1 accuracy in Open Cloze.
A Simple Joint Model for Improved Contextual Neural Lemmatization (N19-1)

Copied to clipboard

Challenge: False positive: a core NLP task of lemmatization seeks to map multiple forms of English verbs to a canonical one, known as the lemma.
Approach: They propose a joint neural model for lemmatization and morphological tagging that achieves state-of-the-art results on 20 languages from the Universal Dependencies corpora.
Outcome: The proposed model achieves state-of-the-art results on 20 languages from the Universal Dependencies corpora.
SCITAT: A Question Answering Benchmark for Scientific Tables and Text Covering Diverse Reasoning Types (2025.findings-acl)

Copied to clipboard

Challenge: Existing scientific question answering datasets lack diverse reasoning types and neglect relevance between tables and text.
Approach: They propose a scientific question answering benchmark for scientific tables and text with diverse reasoning types (SCITAT) to address these challenges, they propose QA benchmark which incorporates tables and texts to ensure that the questions encompass both tables and textes.
Outcome: The proposed benchmark improves by 4.1% over baselines on SCITAT.
TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved impressive results in Machine Translation (MT). human evaluations reveal that LLM-generated translations still contain various errors.
Approach: They propose a LLM-based self-refinement framework that feeds error information back into LLMs to facilitate self-finement, leading to enhanced translation quality.
Outcome: The proposed framework outperforms internal refinement and feedback methods while ensuring a robust translation quality baseline.
Test-time Augmentation for Factual Probing (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve factual probing are relation-specific and do not generalize to unseen relation types.
Approach: They propose to use test-time augmentation to augment and ensemble prompts at test time to reduce sensitivity to prompt variations.
Outcome: The proposed method improves model confidence, but for other models, it leads to degradation.
I Could’ve Asked That: Reformulating Unanswerable Questions (2024.emnlp-main)

Copied to clipboard

Challenge: Existing large language models do not assist users in reformulating unanswerable questions . a recent study found that the models failed to reformulate questions based on assumptions that conflict with or cannot be verified with the information available in documents.
Approach: They evaluate open-source and proprietary LLMs on couldAsk to evaluate their performance . they found that GPT-4 and Llama2-7B successfully reformulate questions only 26% and 12% of the time .
Outcome: The proposed model successfully reformulates questions only 26% and 12% of the time . the proposed model is not able to reformulate questions, but it can be improved .
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints (2026.acl-long)

Copied to clipboard

Challenge: Existing LLM planning benchmarks emphasize local, step-level reasoning rather than global constrained optimization.
Approach: They propose a benchmark for practical long-horizon agent planning that uses local constrained reasoning and global constrained optimization.
Outcome: The proposed benchmarks show that even frontier agentic LLMs struggle with these problems.
Ab Antiquo: Neural Proto-language Reconstruction (2021.naacl-main)

Copied to clipboard

Challenge: Historical linguists have identified regularities in the process of historic sound change.
Approach: They propose a method to reconstruct proto-words based on cognates in daughter languages . they use a dataset of 8,000 comparative entries to analyze phonological changes .
Outcome: The proposed method outperforms conventional methods in a proto-word reconstruction task.
Detecting Subevents using Discourse and Narrative Features (P19-1)

Copied to clipboard

Challenge: Existing models for detecting events as subevents have been developed for analyzing textual understanding.
Approach: They propose a supervised model that automatically identifies when one event is a subevent of another.
Outcome: The proposed model outperforms previous systems on two annotated corpora with event hierarchies, achieving 0.74 BLANC F1 on the Intelligence Community corpus and 0.70 F1 for the HiEve corpus, respectively a 15 and 5 percentage point improvement over previous models.
Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Using a Llama-3.2-3B-Instruct backbone, we evaluate four adaptation mechanisms across four benchmarks: Gorilla APIBench, Spider 2.0, WebArena, and InterCode.
Approach: They compare hypernetwork-based LoRA adaptation against carefully designed few-shot prompting in a controlled experiment . they find that few- shot prompting contributes +21.5% to performance and documentation contributes 0% .
Outcome: The hypernetwork-based LoRA adaptation provides no measurable improvement over few-shot prompting alone.
CBOW-tag: a Modified CBOW Algorithm for Generating Embedding Models from Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Using word2vec, we train distributional semantic models that predict a word from the context or vice versa.
Approach: They propose a modified version of the CBOW algorithm implemented in the fastText framework that includes the representation of original word forms and their annotation at the same time.
Outcome: The proposed model can answer questions such as What do we eat?, What can we do with a skeleton?, etc.
ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on large language model-based agents focus on evaluation benchmarks without training support.
Approach: They propose a large-scale Chinese shopping simulation environment that uses large language models to train agents.
Outcome: The proposed model performs poorly in a large-scale and challenging shopping environment in China.
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction (2026.acl-long)

Copied to clipboard

Challenge: End-to-end (E2E) spoken dialogue systems are replacing cascaded pipelines for voice-based human-AI interaction. Existing benchmarks evaluate these systems on synthetic speech and single-turn tasks, leaving multi-turn conversational ability underexplored.
Approach: They propose an open-source benchmark to evaluate spoken dialogue systems under natural multi-turn interaction patterns.
Outcome: The proposed model fails on the highest-performing model with 54.65% pass rate.
Mind’s Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of multimodal large language models (MLLMs) have demonstrated compelling visual understanding in recent years.
Approach: They propose a multimodal large language model with eight visuo-cognitive tasks inspired by classic human intelligence tests organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation.
Outcome: The proposed frameworks are based on eight visuo-cognitive tasks inspired by human intelligence tests and organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations